Papers with region-level retrieval
OLIVE: Object Level In-Context Visual Embeddings (2024.acl-long)
Copied to clipboard
| Challenge: | Existing vision-language models lack fine-grained object-level understanding and grounding . existing models implicitly align text tokens with image patch tokens, which is ineffective for embedding alignment at the same granularity and introduces noisy spurious background features. |
| Approach: | They propose a method to prompt large language models with in-context visual object vectors . this method allows for controllable object-level reasoning . |
| Outcome: | The proposed method achieves competitive referring object classification and captioning performance while offering zero-shot generalization and robustness to visually challenging contexts. |